Frontiers in Digital Health
○ Frontiers Media SA
Preprints posted in the last 7 days, ranked by how well they match Frontiers in Digital Health's content profile, based on 24 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit.
Langbaum, J. B.; Erickson, C. M.; Langlois, C.; Wood, E. M.; Egleston, B. L.; Harkins, K.; Mim, R.; John, S.; Brown, C.; Brown, S.; Howe, S.; Cacioppo, C.; Eppelmann, L.; Enos, J.; Salata, H.; DeSantiago, D.; Largent, E. A.; Reiman, E. M.; Denkinger, M. N.; Ashton, N. J.; Roberts, J. S.; Karlawish, J.; Bradbury, A. R.
Show abstract
Importance: Patients are increasingly learning Alzheimers disease (AD) genetic and biomarker results through electronic health portals. Evaluation of alternative scalable delivery models for return of AD risk information is needed to best support patient understanding and psychological well-being. Objective: To determine whether a patient-centered digital platform is comparable to clinician-mediated telehealth sessions for returning APOE and plasma pTau-217 results on outcomes of knowledge and psychological well-being. Design: The Evaluation of Self-Mediated Alternatives for Risk Testing Education and Return of Results (eSMARTER) study was a noninferiority trial of a patient-centered digital platform compared to clinician-mediated disclosure of APOE genotype and optional pTau-217 disclosure. Setting: Decentralized, fully remote trial enrolled participants in the contiguous United States (U.S.) between October 2024 and February 2025, with follow-up completed in November 2025. Participants: Eligible participants were aged 60-80 and had previously undergone APOE genotyping (without disclosure) via the GeneMatch program, passed psychological screening, had internet access, and were English-speaking. Interventions: Participants were randomized, 2:1, to the eSMARTER digital platform or clinician-mediated disclosure of APOE genotype. Following the 6-month post-APOE assessment, participants were offered optional pTau-217 disclosure via the same randomized modality. Main Outcomes and Measures: Primary outcomes at 1-7 days following APOE disclosure included changes in anxiety, disease-specific distress, and AD-related knowledge within a priori non-inferiority margins. Results: 674 persons (mean [SD] age 68 [4.7] years; 451 [67%] female; mean [SD] telephone MoCA=19 [2]) were eligible and provided demographic information. 651 participants were randomized to clinician-mediated (n=216) or digital disclosure (n=435) and completed APOE disclosure (66 [10%] APOE4 homozygotes, 377 [58%] heterozygotes, 208 [32%] non-carriers). 604 participants completed the study; 500 completed optional pTau-217 disclosure. Baseline characteristics were balanced across groups. At 1-7 days following APOE disclosure, scores on AD-related knowledge, PROMIS Anxiety, and disease-specific distress measures met non-inferiority. Conclusions and Relevance: Disclosure of APOE genotype by the eSMARTER digital platform is non-inferior to clinician-mediated telehealth disclosure. No significant between group differences were found following disclosure of pTau-217 results. Together, these results suggest that this digital platform may provide an evidence-based scalable approach for returning AD genetic and biomarker results.
Shen, H.; Agorinya, I. A.; Ayanore, M. A.; Brede, M.; Chapman, A.; Head, M.
Show abstract
Introduction Safe and timely blood availability remains a major global health challenge, especially in low- and middle-income countries. Digital tools may accelerate donor contact, but digital reachability alone does not ensure that people will notice, trust and act on urgent requests to support blood donation efforts. We examined factors associated with anticipated engagement in digitally coordinated urgent blood-donor mobilisation among digitally reachable adults in Ghana. Methods We conducted a cross-sectional online survey from September 2025 to January 2026 across Ghana's 16 regions. Participants were recruited via Facebook advertising and snowball sampling. Factors associated with urgent blood-donor mobilisability were assessed under four criteria: high future-donation willingness; high willingness to install a trusted donation app; high willingness to respond to a trusted urgent-request; and high practical flexibility to leave current activities. Descriptive analyses and multivariable logistic regression examined prevalence and associated factors. Results Among 1,067 participants, 577 (54.1%) met all four criteria. Future-donation willingness (91.8%), trusted-app installation willingness (83.2%) and trusted-request response willingness (82.7%) were common, whereas practical flexibility was lower (66.6%). In the adjusted model, high formal health-system trust (adjusted OR (AOR) 3.95, 95% CI 2.08-7.50), high digital-response readiness (AOR 2.26, 1.66-3.08), previous donation (AOR 1.47, 1.08-2.01), high donation knowledge (AOR 1.42, 1.03-1.97) and willingness to donate to strangers were positively associated with high mobilisability. Women (AOR 0.60, 0.43-0.83), participants reporting a work-schedule barrier (AOR 0.43, 0.29-0.66) and those travelling over 30 min to the nearest healthcare facility at night (AOR 0.66, 0.45-0.96) had lower adjusted odds. Conclusions Digital reachability and stated donation willingness may overestimate the population pool available for emergency donation. Digital blood-donor solutions should consider verifiable health-system requests, account for response readiness and current availability, and connect willing individuals with accessible collection options and transport support where needed.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.
Show abstract
The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
Farzana, S.; Arian, A.; Rundek, T.; Desvarieux, M.; Ahsan, H.
Show abstract
Early identification of Alzheimer's disease and related dementias (ADRD) remains challenging despite its importance for timely intervention, management of modifiable risk factors, and care planning. We developed and evaluated ADRD onset prediction models using longitudinal electronic health records (EHRs) from the All of Us Research Program at clinically meaningful lead times of 6, 12, 24, and 36 months before diagnosis, benchmarking interpretable count-based representations against four publicly available pretrained clinical foundation models (CLMBR-T, GPT-style, LLaMA-style, and Mamba) across multiple ADRD phenotype definitions. Count-based models consistently achieved the highest discrimination and calibration across all cohorts and prediction horizons. Predictive performance declined with increasing lead time for all approaches; however, the performance gap between count-based and pretrained representations progressively narrowed, with foundation models achieving comparable AUROC of 0.719 (compared to the AUROC of 0.738 of count-based model) at the 36-month horizon while providing higher sensitivity and F1 scores under a fixed operating threshold. External validation with zero-shot evaluation on UChicago EHRs exhibited limited generalizability for count-based and pretrained clinical foundation model based representations. These findings demonstrate that transparent count-based EHR representations remain the strongest overall approach for ADRD onset prediction, while pretrained clinical foundation models provide complementary advantages for long-term risk identification and establish a benchmark for evaluating transferable clinical representations in temporal ADRD risk prediction.
Okundaye, D. O.; Isiekwene, C. C.
Show abstract
Acute kidney injury (AKI) is a frequent complication within intensive care units, with its sudden onset often missed. This is especially important because a timely window for intervention is required as delayed detection leads to progressively worse outcomes. Existing machine learning and deep learning models have contributed to closing this gap, but their complexity, requiring hundreds to thousands of features, and lack of generalisation pose a limitation that prevents them from being integrated into clinical workflows across different electronic health-record ecosystems. This study presents a 37-feature XGBoost model trained on the MIMIC-IV dataset with 5.4% positive cases, with hyperparameters optimised via Optuna and probabilities calibrated using isotonic regression, designed for transportability across clinical settings. Validation was conducted internally using a temporal patient-level split simulating prospective deployment, training on 2008-2016 data and testing on 2017-2022 data"External validation was performed on the eICU Collaborative Research Database, a multi-centre dataset spanning 208 US hospitals, using the trained model without retraining. SHAP TreeExplainer was used to provide feature-level explainability for individual predictions. Internal testing yielded an AUROC score of 0.794 for predicting AKI onset within a 12-24 hour window. External validation produced a 0.750 AUROC without retraining. Equitable discrimination was observed across gender, age, chronic kidney disease presence, race, and AKI stages on both datasets, with a 95% internal CI of 0.789-0.799 confirming the model's estimate stability. These results suggest that clinically useful prediction systems are achievable with substantially fewer features than current models require.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Knol, L.; Nagpal, A.; Hussain, F.; Beckmann, C. F.; Leow, A.; Eisenlohr-Moul, T. A.; Marquand, A. F.
Show abstract
Digital phenotyping, which is defined as quantifying someone's behaviour with digital devices, provides unprecedented opportunities for understanding human mental health but is hampered by high levels of inter-individual variability. Here, we propose a new method to address this, parsing inter-individual variability by decomposing the digital phenotype dynamics into latent trajectories and using each individual's trajectory membership as a moderator when modelling psychopathology over the same timeframe. We applied our method in the context of mood symptom exacerbation across the menstrual cycle, where symptom severity and timing are inconsistent between individuals. Using the BiAffect platform to collect smartphone typing dynamics, we found stable trajectories in smartphone movement rate: one group of participants showed substantial movement rate fluctuations across the menstrual cycle, whilst the others did not. Participants with movement fluctuations displayed increased fluctuations across the cycle in prospective anhedonia and depression ratings, but not in anxiety, irritability, and suicidal ideation.
Xiang, S.; He, H.; Xie, Z.; Cheng, C.-Y.; Li, H.; Liu, D.
Show abstract
Agentic workflows can coordinate modelling, but balancing predictive performance, measurement burden and reproducibility is unclear. We developed DXA Agent, an agentic workflow for dual-energy X-ray absorptiometry (DXA) outcomes integrating planning, feature-model refinement, tools, provenance and hypothesis-generating interpretation. Models were independently developed and tested in UK Biobank (5,318 participants) and the National Health and Nutrition Examination Survey (NHANES; 3,777 participants), using cost-efficient and no-limit strategies. Across 20 UK Biobank and three NHANES bone mineral density sites, cost-efficient models achieved lower RMSE and higher R2 than the best conventional comparator, with median relative RMSE reductions of 10.9% and 9.9%, respectively. Classification was task dependent: UK Biobank osteoporosis averaged AUROC 0.839 and PR-AUC 0.182, whereas NHANES performance was comparable with conventional models. Higher-burden features did not consistently improve prediction. These retrospective, cohort-internal findings position DXA Agent as an inspectable, measurement-burden-aware research workflow requiring independent prospective validation.
de Araujo Morais, J. H.; Dias Ferreira, C.; Saraceni, V.; Medeiros de Oliveira Cruz, D.; Mateus Oliveira Aguilar, G.; Cruz, O. G.
Show abstract
Motivation: With the scaling frequency and intensity of extreme heat events across the globe, it is critical for public institutions to develop early detection systems and continuous monitoring of these events and their impacts. In Brazil, Rio de Janeiro was the first city to publish its heat protocol, with the Rio Heat Dashboard as a central component of this system. Implementation: The dashboard was implemented using R/Shiny and integrates climatic and health data from multiple sources. General features: The application comprises real-time heat exposure monitoring and automatic alert level classification, which is monitored daily by multiple municipal actors and supports activation of actions specified in the heat protocol. It also features a health impact module, which lists each heat event and its impact on mortality, and primary care and emergency visits. Availability: The source for full reproducibility is available through https://github.com/joaohmorais/RioHeatDashboard.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.
ye, y.; Zeng, Z.; Tian, X.; Yuan, Z.; Wang, J.; Zhu, Y.
Show abstract
Artificial intelligence applied to routine electrocardiograms (ECGs) has largely focused on detecting existing disease or predicting individual cardiovascular outcomes. Whether ECGs can support prediction of multiple future diseases across organ systems remains unclear. We developed ECG-RISK, a multitask survival model for 67 incident three-character ICD-10 endpoints using ECG waveforms, demographic characteristics and routinely collected laboratory data from 86,673 MIMIC-IV patients. Discrimination was highest for heart, brain, kidney and lung endpoints, with organ-level C-indices ranging from 0.796 to 0.825, whereas liver and pancreatic endpoints showed lower discrimination. The ECG-only model achieved strong discrimination across most endpoints, whereas the incremental improvement gained by incorporating ECG and laboratory inputs beyond demographic information varied substantially across endpoints. Across the nine exploratory aggregated outcomes, Kaplan Meier curves showed clear separation among model-score tertiles. Discrimination was highest for dementia (C-index, 0.891) and heart failure (C-index, 0.857). These findings support the feasibility of ECG-based longitudinal risk prediction across multiple diseases. External validation and competing-risk analyses are required to assess generalisability and clinical utility.
Lu, Z.; Uddin, S.; Uribe, S.; White, S.; Martins, R. T.; Chau, S.; Mosaddek, A. S. M.; Islam, M. S.; Nahar, N.; Azad, A. K. M.; Hossain, K. M. N.; Choudhury, H. S.; Hasan, K. M. R.; Mosaddek, N.; Rahman, S.; Hossain, M. M.; Sizar, K. M. M. H.; Angione, C.; Lio, P.; Islam, M. T.; Moni, M. A.
Show abstract
Stroke remains a leading cause of mortality and long-term disability worldwide, yet rapid diagnosis is often limited by the shortage of trained radiologists, particularly in resource-constrained settings. Automated analysis of CT imaging offers a potential solution, but existing methods often struggle to achieve clinically generalisable performance while jointly addressing multiple diagnostic tasks. Here we present the Intelligent Integrated Stroke Diagnosis System IISDS, an end-to-end deep learning framework built upon StrokeGNN, a graph-based architecture that integrates 3D contextual feature extraction with U-Net-based 2D lesion segmentation to enable comprehensive stroke analysis from non-contrast CT scans. IISDS performs stroke subtype classification, lesion segmentation and lesion volume estimation within a unified pipeline. To develop and validate the system, we collected and curated BGD-ISD through a collaboration between AI researchers, neurologists, radiologists and clinicians, resulting in a large multi-centre dataset comprising 1,507 CT scans from 597 stroke cases acquired across six hospitals and medical centres in Bangladesh. Across BGD-ISD and multiple publicly available datasets, IISDS achieves state-of-the-art performance on all tasks, improving segmentation accuracy by [≥]0.011 Dice score, reducing lesion volume estimation error by [≥]0.3 average symmetric surface distance (ASSD), and increasing classification performance by [≥]0.018 area under the receiver operating characteristic curve (AUC) compared with existing approaches. These results demonstrate the potential of graph-based deep learning to enable clinically generalisable, automated and scalable stroke diagnosis from CT imaging, supporting rapid clinical decision-making, particularly in healthcare environments with limited access to expert radiological interpretation.
Kremer, P.; Schlicker, N.; Hasnaj, R.; Bamberger, J.; Witte, T.; Haase, I.; Mayr, A.; Schmidt, C.; Osteras, N.; Baraliakos, X.; Kuhn, S.; Krusche, M.; Knitza, J.
Show abstract
Objectives To evaluate whether access to a certified large language model (LLM)-based clinical decision support system improves physician diagnostic performance in rheumatology compared with conventional diagnostic resources alone. Methods In this multicentre, open-label, randomised controlled trial, 82 physicians from seven hospitals in two countries were randomised 1:1 to conventional diagnostic resources plus Prof. Valmed or conventional resources alone. Participants assessed three rheumatology vignettes before and after assistance. The primary outcome was top-1 diagnostic accuracy. Secondary outcomes included top-3 accuracy, diagnostic reasoning, confidence, case-processing time and perceived support quality. Results Top-1 accuracy increased from 22.2% to 33.3% in the intervention group and from 23.3% to 35.0% in the control group, with no between-group difference in improvement (adjusted OR 0.99, 95% CI 0.45 to 2.19; p=0.979). Differences in top-3 accuracy, diagnostic reasoning and confidence were also not significant. Assisted case-processing time was substantially shorter with LLM support (94 vs 206 s; adjusted mean difference -112 s, 95% CI -141 to -83; p<0.001). Information timeliness and perceived diagnostic support quality were rated significantly higher in the intervention group. Exploratory analyses showed persistent overconfidence and substantial AI over-reliance. Conclusions Certified LLM-based diagnostic support did not improve diagnostic accuracy compared with conventional resources, but substantially reduced case-processing time and improved perceived support quality. These findings suggest potential workflow benefits while highlighting overconfidence and over-reliance as important safety considerations.
Radoynova, M.; Benouis, M.; schulze, f.; Winter, S.; Bornhauser, M.; Middeke, J. M.; Eckardt, J.-N.
Show abstract
Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.
Lebmeier, A.; Lindner, T.; Karl, C.; Schöler, T.; Rank, A.
Show abstract
Background: Immunochemotherapy (ICT) is considered standard in regards to care for small-cell lung cancer (SCLC) in extensive stages, yet reliable biomarkers for treatment response remain elusive. While previous univariate analyses suggest specific peripheral lymphocyte subsets correlate with survival, the systemic immune response involves complex, multivariate interactions that require advanced analytical approaches. Methods: This paper analysed high-dimensional flow cytometry data from 32 patients with stage IV SCLC treated with carboplatin, etoposide, and atezolizumab. Peripheral blood was analysed at baseline (V0) and longitudinally during treatment. To identify potential early predictive biomarkers and mitigate sample attrition in later cycles, we focused on baseline and measurements after two cycles of ICT (V1). We employed a rigorous machine learning framework utilising nested cross-validation, bootstrapping, and permutation-based statistical testing to evaluate eleven different regression and survival models. Results: Under model-appropriate metrics, regressors did not generalise (R2 <0); conversely, censoring-aware Random Survival Forests (RSF) successfully extracted robust prognostic signatures. Baseline immune profiles (V0) achieved a concordance index (C-index) of 0.66 (p= 0.015), while dynamic changes from V0 to V1 ({triangleup}V) achieved a C-index of 0.65 (p= 0.022). Crucially, absolute values measured after two cycles of ICT (V1) yielded no significant signal (p= 0.445). Feature importance analysis confirmed the prognostic value of Th17 normalisation and identified Naive Regulatory T cells and Memory B cells as candidate components. Conclusion: Machine learning validation confirms a predictive signal in the peripheral immune profile of SCLC patients. Early dynamic shifts in the balance between regulatory and effector immune arms are associated with prognosis, contrasting with the lack of signal in absolute counts after two cycles of ICT. These findings establish a proof of concept for multivariate liquid biopsy immune profiling, warranting confirmation in larger cohorts and highlighting the necessity of integrating systemic and tumour-intrinsic data.
Quigley, H.; Gardiner, B.; McDaid, L.; O'Donnell, C.
Show abstract
Autism Spectrum Disorder (ASD) is a heterogeneous neurodevelopmental condition defined by differences in social communication and restricted, repetitive behaviours. As diagnostic criteria have broadened, ASD is now recognised across a wider range of individuals, raising key questions about its structure: does ASD have discrete sub-types, or is it better conceptualised as a continuous, possibly multidimensional, condition? We aim to explore whether a multidimensional continuum model more accurately captures the variability within ASD. We analysed a large SPARK phenotypic dataset of medical history and diagnostic surveys (background history, SCQ, RBS-R; n=36,710 individuals). We apply and compare two traditional statistical approaches, Factor Analysis and Gaussian Mixture Models, with a modern machine learning technique, the Variational Autoencoder (VAE). VAEs reconstructed unseen test data with ~4-fold better accuracy than Factor Analysis, and ~8-fold better accuracy than Gaussian Mixture Models. We identified four stable latent factors across 100 independently trained VAEs. These four dimensions provide an individual behavioural profile that can be visualized using radar-plots, offering a compact way to compare profiles at the person level. Through further analysis, we found evidence for 3 overlapping clusters or subtypes of ASD identified within the 4D latent space. This work aims to inform new ways of modelling ASD using a VAE that will be able to discern between a continuum or a clustered output and that go beyond binary diagnosis, instead reflecting the complex range of trait profiles, with implications for personalised diagnosis and intervention.
Ji, J.; Sun, Z.; Ying, X.; Hao, J.; Fu, Z.; Shi, D.; Kong, X.; Xu, Y.; Zhang, X.; Du, X.; Zhang, Z.; Liu, X.; Lin, P.; Wang, H.
Show abstract
Background. Routine service databases are attractive sources of training labels for clinical prediction models, but the processes that write those labels are rarely audited before the labels are used. In a deployed community cognitive-screening programme, we audited the routine cognitive-status label, built a matrix of twenty-four model arms over the same patients under a specialist reference standard, and measured what each supervision choice bought or cost. Methods. The study cohort is the 672 individuals whose cognitive status was recorded by a titled (attending-or-above) physician, that record being the reference standard; after holding out one institution entirely, a development panel of 642 individuals at 38 institutions. The routine cognitive-status label these individuals also carry was first audited at the operator level: for each data-entry account we counted diagnoses entered and the proportion recording any impairment, and tested a competing bulk-timestamp explanation. Twenty-four arms span the supervision choices such a programme faces: an incumbent 21-variable logistic regression; local language models (Qwen2.5-1.5B/3B, Qwen3-4B/8B) zero-shot, with chain-of-thought, fine-tuned on physician labels, on routine labels with and without decontamination, or on a proxy scale-band task; preference-optimised (DPO) and reinforcement-trained (GRPO) variants; a proprietary frontier model queried zero-shot; and knowledge distillation of that frontier model into the regression and into the local 4B, using 943 teacher-labelled records from the programme's unlabelled pool. All arms are scored out-of-fold under one five-fold split grouped on registry-resolved institution clusters (no cluster spans a fold); paired contrasts use a 2,000-draw cluster bootstrap. Results. 181 operator accounts (each entering at least 100 diagnoses with zero recorded impairments) account for 45,315 rows - 40.5% of the outcome column; recorded impairment falls monotonically with account volume (15.7% for 1-9 rows to 0.7% for 500-999); a bulk-timestamp explanation was tested and refuted, identifying the write-time column as a migration artefact. Under the specialist standard, no locally fine-tuned arm beat the incumbent regression (AUROC 0.926): physician-label SFT reached 0.924 (4B), DPO 0.881, and GRPO 0.789; the pre-registered two-stage proxy-then-RL recipe was worse than its single-stage contaminated baseline (-0.030, 95% CI -0.077 to -0.004). Chain-of-thought reduced discrimination at every size (-0.072, -0.080, -0.041 at 1.5B/3B/4B; -0.012, n.s., at 8B). The frontier model scored 0.932 (vs. regression +0.007, n.s.). The distilled 4B reached 0.940 - above the incumbent (+0.014, 0.004 to 0.031) and above its own teacher (+0.008, 0.001 to 0.017) - with near-teacher calibration; it reached the teacher's level by 50 teacher labels and changed little beyond 200. Conclusions. The audit and the arm matrix support one deployment recipe: audit the routine label at the operator level before training on it; do not expect fine-tuning, preference optimisation, or reinforcement learning on a few hundred specialist cases to beat a well-calibrated regression; and if a frontier model is available but undeployable, spend a bounded number of queries on it as a labelling instrument and distil. A companion paper uses these frozen predictions to quantify how evaluation design choices compare with model choice.